↓Skip to main content
  1. Showcases/

Case study - CEREBRA, an integrated data engineering platform for brain science research

CEREBRA demonstrates how ABCD’J’s open, interoperable technologies can be combined into a self-hosted research data engineering platform for collaborative brain science.

Research collaborations need infrastructure not only to store data, but also to manage the metadata, provenance, access permissions, and workflows that make data usable and reusable. With collaborative research projects being distributed across groups and institutions, this infrastructure must also connect different tools and systems while accommodating the requirements of individual projects and scientific applications.

Within the ABCD-J context, CEREBRA provides a concrete deployment of this approach for brain research conducted by Forschungszentrum Jülich and its partners. Operated by the Institute for Neurosciences and Medicine - Brain and Behavior (INM-7) and hosted on the Jülich Supercomputing Centre Cloud, CEREBRA brings together a set of open-source technologies from the ABCD-J software stack for collaborative (meta)data management. It provides a modular technical foundation that can be adapted to different research communities, data models, access requirements, and workflows.

The platform #

CEREBRA (https://cerebra.fz-juelich.de/) combines several interoperable components that address different aspects of data engineering:

  • Forgejo-aneksajo provides collaborative infrastructure for managing code and data, extending Forgejo with support for large files through git-annex.
  • DataLad and git-annex support large-scale, versioned data management and reproducible access to datasets.
  • Dump Things Service provides structured metadata storage and backend APIs.
  • shacl-vue provides user-facing interfaces for entering and retrieving metadata based on shared schemas.

Together, these components provide a common technical stack that can be deployed and configured for different research groups and use cases.

Connecting data and metadata #

A data engineering platform needs to connect two related questions: What research data exist, and how can those data actually be accessed and used? Rather than treating (meta)data management and data discovery as separate systems, CEREBRA demonstrates how these capabilities can be connected into an integrated research infrastructure.

Shared metadata models can define what information is collected and how it is structured. This information can then be mapped to broader semantic models, providing a basis for connecting local data infrastructures to external discovery and knowledge systems, including the Helmholtz Knowledge Graph.

Domain-specific services can similarly be integrated into the platform to extend the ways in which datasets can be discovered and used. For brain research, this includes connecting CEREBRA with services such as Neurobagel and EBRAINS.

The significance #

CEREBRA illustrates several principles at the heart of the ABCD-J platform:

  • Open technologies allow research communities to build on and extend existing software rather than depending on proprietary platforms.
  • Self-hosted infrastructure gives communities greater control over their data, services, and workflows.
  • Shared metadata models provide a common language for describing research resources while remaining adaptable to individual use cases.
  • Reusable components allow the same technical building blocks to be deployed across different projects and institutions.
  • Interoperability connects local research infrastructure with other tools, services, and knowledge systems.

An important aspect of CEREBRA is that it is not intended to be limited to a single research group or a single configuration. The current and planned user base includes research groups within the Forschungszentrum Jülich (FZJ) Institute of Neurosciences and Medicine (INM) as well as collaborative research initiatives such as TRR379 and SFB1451, alongside other FZJ research activities. The same underlying stack can support these different communities and their local requirements while avoiding the need to develop and maintain a separate technical infrastructure for each use case.

This approach provides a pathway toward broader deployment across Forschungszentrum Jülich and the Helmholtz Association, and potentially to external research environments. In particular, CEREBRA can serve as a technical component in collaboration with surrounding university medical centers, providing a foundation for connecting ABCD-J research infrastructure with clinical and research partners. This creates a bridge between the current brain-research deployment and future data engineering activities aimed at supported distributed research communities across groups, institutions, and scientific domains.

There's no articles to list here yet.